Papers by Yavuz Faruk Bakman
Do Not Design, Learn: A Trainable Scoring Function for Uncertainty Estimation in Generative LLMs (2025.findings-naacl)
Copied to clipboard
Duygu Nur Yaldiz, Yavuz Faruk Bakman, Baturalp Buyukates, Chenyang Tao, Anil Ramakrishna, Dimitrios Dimitriadis, Jieyu Zhao, Salman Avestimehr
| Challenge: | Existing methods for probability-based UE are limited by their inability to handle biased probabilities and complex semantic dependencies between tokens. |
| Approach: | They propose a learning-based scoring function that captures complex dependencies between tokens and probabilities and produces more reliable responses. |
| Outcome: | The proposed function outperforms existing scoring functions in question-answering and arithmetical reasoning tasks with different datasets. |
TruthTorchLM: A Comprehensive Library for Predicting Truthfulness in LLM Outputs (2025.emnlp-demos)
Copied to clipboard
Duygu Nur Yaldiz, Yavuz Faruk Bakman, Sungmin Kang, Alperen Öziş, Hayrettin Eren Yildiz, Mitash Ashish Shah, Zhiqi Huang, Anoop Kumar, Alfy Samuel, Daben Liu, Sai Praneeth Karimireddy, Salman Avestimehr
| Challenge: | Generative Large Language Models (LLMs) produce untruthful outputs, referred to as hallucinations, which are often referred as false positives. |
| Approach: | They propose an open-source Python library with over 30 truthfulness prediction methods. |
| Outcome: | The proposed methods span diverse trade-offs in computational cost, access level, grounding document requirements, and supervision type (self-supervised or supervised). |
Reconsidering LLM Uncertainty Estimation Methods in the Wild (2025.acl-long)
Copied to clipboard
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Sungmin Kang, Tuo Zhang, Baturalp Buyukates, Salman Avestimehr, Sai Praneeth Karimireddy
| Challenge: | Existing studies evaluate UE methods in short-form QA settings, but real-world deployment presents several challenges. |
| Approach: | They examine UE methods' sensitivity to decision threshold selection and their robustness to query transformations such as typos and adversarial prompts. |
| Outcome: | The proposed methods exhibit robustness against typos, adversarial prompts, and prior chat history, and are highly susceptible to adversarials. |
Un-considering Contextual Information: Assessing LLMs’ Understanding of Indexical Elements (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in coreference resolution tasks, but previous studies only assessed performance with nouns and third person pronouns. |
| Approach: | They evaluate LLMs' performance on coreference resolution with indexicals like I, you, here and tomorrow which come with unique challenges due to their linguistic properties. |
| Outcome: | The proposed models perform well with some indexicals while struggling with others. |
MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs (2024.acl-long)
Copied to clipboard
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr
| Challenge: | Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. however, their tendency to produce inaccurate or misleading outputs poses a potential risk. |
| Approach: | They propose a new scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. |
| Outcome: | The proposed scoring function improves UE performance on a medical QA dataset. |